How do I compare two images to see if they are equal?
Often people ask how can they compare images (probably meaning the contents of images) to see whether they are equal, hoping for a simple function or operator that returns true or a value between 0 and 1 where 0 means different, 1 means equal and intermediate values can be considered different degrees of similarity. Many applications would benefit of this kind of high-level comparison, for example, for detection of image copyright violations, face identification in crowds for security purposes, semantic image searching, visual navigation for robots, etc.
Unfortunately comparison of images contents is not a simple task, and in most cases, a ill-defined one.
For example, consider the images below (move the mouse pointer over the images to see a description). Most of those depict the same animal, so those could be considered equal for some purposes (e.g. while searching the web for mammals behind fences), or completely different for other purposes (e.g. looking for squirrels facing up or large mammals behind fences).
![]() A |
![]() B |
![]() C |
![]() D |
![]() E |
![]() F |
Similarity examples
Most people think about the contents of the image without realizing how complex its processing by the human visual system and brain is. Do an experiment, show the first image to different people and ask "what's in this image?" -- probably most will describe the animal in the image (since it can be perceived as the subject in it), but the description itself will probably vary in completeness (number of relevant objects perceived in the image) and detail (amount of relevant information about each object). If not even humans can come up with a simple, concise description of what's in an image (so we could compare the descriptions of the contents), how can we expect a computer to do that with a simple algorithm?
The problem is that computers cannot "see" an image and "compare" it to other to determine if they are equal, where "equal" refers to the images' contents. Granted, advances in image processing and computer vision may one day make this task feasible under controlled conditions, but not now, and definitely not with simple algorithms. Consider how many variables we have to deal with: each object on the images have a shape, size, orientation, color (maybe more than one!), texture, etc.; the same object may appear different in images, e.g. it can appear in a different position, scale, orientation or illumination conditions or combinations of those; it can be partially ocluded; should we consider these variations as irrelevant for the comparison or are they important? Should we assign different "weights" for the objects for similarity comparison so similarity measures in some are more important than in others? Small objects can be missing in one of the images; should we ignore them? Is "background" an object and should it be considered in the comparison, and what, exactly, is the background?
Comparing images in regard to their contents
If you want to compare images semantically, so some images on the example above will be considered equal (or very similar, with exception of 'F'), you will face a hard task that cannot be solved with simple algorithms. You will possibly have to process the image to extract a set of features that can be mapped to semantic objects even if those features are different in regard to some aspects (e.g. position, scale) and compare the objects' features instead of the pixel or image features.
You can see that image contents comparison is not a simple task, but it can be made simpler if we decide to reduce the problem to a simpler one -- instead of asking "are the object in these images equal?" we could ask "are some of the regions in this image similar, in some aspect, to regions in the other image?". Now we can deal with the problem with some image processing (and artificial intelligence) techniques.
Some steps that could solve the problem are:
Note that the steps above are as generic as possible, and does not ensure success for a particular task. Note also that there are many possible algorithms that can be used in some steps, and each of those have many variations and parameters, therefore blind application of a off-the-shelf algorithm will probably lead to failure -- it is very important to understand what you are doing and why so you can expect useful, consistent results.
Note also that the process used to extract semantic objects in the human brain is very, very complex and flexible, being able to map not only the semantic objects, but their generic behaviour on the scenes and even their categories as well. It is very easy for us to point to the images that have a "squirrel facing up" or "small rodent hanging in a fence", but those tasks cannot be done with only image processing algorithms.
Comparing images for possible similarity without considering the high-level contents
In some applications a very superficial similarity measure between images can be useful. For example, consider the task of locating images approximately similar in content to one you have, but without any guarantee whatsoever that they represent or contain the same object -- something similar to a rough, passing similarity. An obvious application would be locate all images in a disk that can be similar to a chosen one.
There are several possible ways to do that, but all of them require the extraction of some feature from the image to comparison with the features from other images and the calculation of a similarity measure. Both the feature extraction and similarity calculation algorithms can be as simple or complex as we want.
As an example, let's choose a simple feature extraction method and similarity measure, both shown below:
|
The features for our test will be 25 RGB triples, corresponding to the average of the RGB values on the 25 regions marked in the figure on the left. The image will be normalized to 300x300 pixels. No texture or variance feature will be stored, only the color averages. Each region has 30x30 pixels. Each image will be represented, then, a 25x3 feature vector. To calculate the similarity measure between two images A and B we will take each of the 25 regions, calculate the Euclidean distance between the regions and accumulate.
The distance from A to A will be, by definition, zero. The upper bound (maximum possible distance between two images, using this similarity measure method) is calculated as 25*(Math.sqrt( (255-0)*(255-0)+ (255-0)*(255-0)+ (255-0)*(255-0) )) or a little bit over 11041. |
Features for similarity
This method was chosen because it is simple to understand and implement and can be easily modified by the reader. It combines color (spectral) information with spatial (position/distribution) information, and is expected to be more robust (i.e. tolerant to differences) than comparing pixel by pixel or the average of the whole image, but, again, it is very simple and cannot be expected for perform well in any circumstances, being shown only as an example.
To test the feature extraction and similarity measure we will use a set of 24 test images, shown below. Some of those images have similar objects on them (wall, trees, sky) but we are not considering meaning on the images, just patches of similar colors. Images are in different sizes, click on the icons to get the full images.
The 16 images on the first two rows are photos of objects, while the last row is of images from the first two rows distorted in scale, color, position, etc.
Similarity test regions
The application NaiveSimilarityFinder.java, show below, has a GUI which allows the user to select a reference image A and then get all images in the same directory, normalize them to the same size (300x300), extract the features, calculate the distance from the features of A and show them in order of less distant to most distant.
The application NaiveSimilarityFinder.java needs an auxiliary class, JPEGImageFileFilter.java, shown below.
Some test runs of the application with that image data set are shown below. The first image is a row is the reference image, and the others image in the row are best seven matches are shown (except for the match for the same image), then the worst two matches.
| Original | Best 7 matches | Worst 2 matches | |||||||
2037.415 |
2065.896 |
2071.409 |
2141.598 |
2377.580 |
2480.289 |
2483.258 |
4627.606 |
4985.659 | |
98.661 |
1922.567 |
2048.512 |
2333.559 |
2455.076 |
2501.762 |
2531.955 |
4397.305 |
5372.516 | |
1470.144 |
2009.162 |
2048.512 |
2088.469 |
2133.800 |
2314.644 |
2378.338 |
4687.977 |
6023.520 | |
176.247 |
2026.587 |
2037.415 |
2062.042 |
2447.092 |
2758.626 |
2759.446 |
5247.339 |
5923.702 | |
354.554 |
1709.610 |
2151.522 |
2406.494 |
2677.916 |
2718.740 |
2755.746 |
4557.786 |
5348.768 | |
719.069 |
1340.927 |
2289.929 |
2328.654 |
2372.341 |
2377.580 |
2468.600 |
3965.858 |
4291.338 | |
344.428 |
1892.389 |
2372.341 |
2602.972 |
2633.854 |
2818.922 |
2826.561 |
5673.467 |
6104.159 | |
Naive similarity comparison results
We can see that some results are more expected or consistent than others, but the algorithm performed quite acceptably, being able to identify, in most cases, the distorted versions when the originals were used as the references, as long as the distortion does not makes the versions very different and the colors are not changed much.
This application could successfully help organize a large collection of digital photographies when there are many repeated shots or when the photographer took many shots of the same subject with small adjustments on the camera (e.g. photos of groups, to make sure no one has blinked).
It should be noted that the application could be modified to calculate the distance from each image to all the others, which could be used then to create a dendogram through a hierarchical clustering algorithm. The dendogram could then be used to separate the images in groups of images which are similar considering the features and measure used.
Many improvements could be done in this simple algorithm: filtering of the regions to reduce influence of noise in the average value, different weights for different regions (e.g. central regions could be considered more important than edges), distance measure in other color space than RGB, etc.
This list is far from complete:
For related information, see A Brief Tutorial on Image Classification and How do I find whether an image contains a specific object?.
![]() |
![]() |